Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Contrastive Language-Image Pre-training</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Contrastive_Language-Image_Pre-training"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.math.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.pygments.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Contrastive_Language-Image_Pre-training rootpage-Contrastive_Language-Image_Pre-training skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Contrastive Language-Image Pre-training</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<style data-mw-deduplicate="TemplateStyles:r1295905060">
/* start https://en.wikipedia.org/ */


.mw-parser-output .infobox-subbox{padding:0;border:none;margin:-3px;width:auto;min-width:100%;font-size:100%;clear:none;float:none;background-color:transparent}.mw-parser-output .infobox-3cols-child{margin:auto}.mw-parser-output .infobox .navbar{font-size:100%}@media screen{html.skin-theme-clientpref-night .mw-parser-output .infobox-full-data:not(.notheme)>div:not(.notheme)[style]{background:#1f1f23!important;color:#f8f9fa}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .infobox-full-data:not(.notheme)>div:not(.notheme)[style]{background:#1f1f23!important;color:#f8f9fa}}@media(min-width:640px){body.skin--responsive .mw-parser-output .infobox-table{display:table!important}body.skin--responsive .mw-parser-output .infobox-table>caption{display:table-caption!important}body.skin--responsive .mw-parser-output .infobox-table>tbody{display:table-row-group}body.skin--responsive .mw-parser-output .infobox-table th,body.skin--responsive .mw-parser-output .infobox-table td{padding-left:inherit;padding-right:inherit}}


/* end https://en.wikipedia.org/ */
</style><table class="infobox vevent"><tbody><tr><th colspan="2" class="infobox-above summary">CLIP</th></tr><tr><td colspan="2" class="infobox-image logo"></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;"><a href="Programmer" title="Programmer">Developer(s)</a></th><td class="infobox-data"><a href="OpenAI" title="OpenAI">OpenAI</a></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;">Initial release</th><td class="infobox-data">January 5, 2021</td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;"><a href="Repository_(version_control)" title="Repository (version control)">Repository</a></th><td class="infobox-data"><span class="url"><a rel="nofollow" class="external text" href="https://github.com/OpenAI/CLIP">github<wbr>.com<wbr>/OpenAI<wbr>/CLIP</a></span></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;">Written in</th><td class="infobox-data"><a href="Python_(programming_language)" title="Python (programming language)">Python</a></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;"><a href="Software_license" title="Software license">License</a></th><td class="infobox-data"><a href="MIT_License" title="MIT License">MIT License</a></td></tr><tr><th scope="row" class="infobox-label" style="white-space: nowrap;">Website</th><td class="infobox-data"><span class="url"><a rel="nofollow" class="external text" href="https://openai.com/research/clip">openai<wbr>.com<wbr>/research<wbr>/clip</a></span></td></tr></tbody></table>
<p><b>Contrastive Language-Image Pre-training (CLIP)</b> is a technique for training a pair of <a href="Artificial_neural_network" class="mw-redirect" title="Artificial neural network">neural network</a> models, one for image understanding and one for text understanding, using a <a href="Contrastive_learning" class="mw-redirect" title="Contrastive learning">contrastive</a> objective.<sup id="cite_ref-:0_1-0" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
This method has enabled broad applications across multiple domains, including cross-modal retrieval,<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> text-to-image generation,<sup id="cite_ref-stable-diffusion-github_3-0" class="reference"><a href="#cite_note-stable-diffusion-github-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> and aesthetic ranking.<sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Algorithm">Algorithm</h2></div>

<p>The CLIP method trains a pair of models contrastively.<sup id="cite_ref-:0_1-1" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> One model takes in a piece of text as input and outputs a single vector representing its semantic content. The other model takes in an image and similarly outputs a single vector representing its visual content. The models are trained so that the vectors corresponding to semantically similar text-image pairs are close together in the shared vector space, while those corresponding to dissimilar pairs are far apart.<sup id="cite_ref-:0_1-2" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p>To train a pair of CLIP models, one would start by preparing a large dataset of image-caption pairs. During training, the models are presented with batches of <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle N}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>N</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle N}</annotation>
</semantics>
</math></span><img src="./f5e3890c981ae85503089652feb48b191b57aae3.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:2.064ex; height:2.176ex;" alt="{\displaystyle N}" loading="lazy"></span> image-caption pairs. Let the outputs from the text and image models be respectively <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle v_{1},...,v_{N},w_{1},...,w_{N}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>.</mo>
<mo>.</mo>
<mo>.</mo>
<mo>,</mo>
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>N</mi>
</mrow>
</msub>
<mo>,</mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mn>1</mn>
</mrow>
</msub>
<mo>,</mo>
<mo>.</mo>
<mo>.</mo>
<mo>.</mo>
<mo>,</mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>N</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle v_{1},...,v_{N},w_{1},...,w_{N}}</annotation>
</semantics>
</math></span><img src="./182a9d6a14eccf955f18bf86696758fd0bd849ec.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:22.448ex; height:2.009ex;" alt="{\displaystyle v_{1},...,v_{N},w_{1},...,w_{N}}" loading="lazy"></span>. Two vectors are considered "similar" if their dot product is large.
</p><p>The loss incurred on this batch is the multi-class N-pair loss,<sup id="cite_ref-5" class="reference"><a href="#cite_note-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> which is a symmetric <a href="Cross_entropy_loss" class="mw-redirect" title="Cross entropy loss">cross-entropy loss</a> over similarity scores:<span class="mwe-math-element mwe-math-element-block"><span class="mwe-math-mathml-display mwe-math-mathml-a11y" style="display: none;"><math display="block" xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle -{\frac {1}{N}}\sum _{i}\ln {\frac {e^{v_{i}\cdot w_{i}/T}}{\sum _{j}e^{v_{i}\cdot w_{j}/T}}}-{\frac {1}{N}}\sum _{j}\ln {\frac {e^{v_{j}\cdot w_{j}/T}}{\sum _{i}e^{v_{i}\cdot w_{j}/T}}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo>−<!-- − --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mn>1</mn>
<mi>N</mi>
</mfrac>
</mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</munder>
<mi>ln</mi>
<mo>⁡<!-- ⁡ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mo>/</mo>
</mrow>
<mi>T</mi>
</mrow>
</msup>
<mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</munder>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mo>/</mo>
</mrow>
<mi>T</mi>
</mrow>
</msup>
</mrow>
</mfrac>
</mrow>
<mo>−<!-- − --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mn>1</mn>
<mi>N</mi>
</mfrac>
</mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</munder>
<mi>ln</mi>
<mo>⁡<!-- ⁡ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mo>/</mo>
</mrow>
<mi>T</mi>
</mrow>
</msup>
<mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</munder>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mrow class="MJX-TeXAtom-ORD">
<mo>/</mo>
</mrow>
<mi>T</mi>
</mrow>
</msup>
</mrow>
</mfrac>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle -{\frac {1}{N}}\sum _{i}\ln {\frac {e^{v_{i}\cdot w_{i}/T}}{\sum _{j}e^{v_{i}\cdot w_{j}/T}}}-{\frac {1}{N}}\sum _{j}\ln {\frac {e^{v_{j}\cdot w_{j}/T}}{\sum _{i}e^{v_{i}\cdot w_{j}/T}}}}</annotation>
</semantics>
</math></span></span>In essence, this loss function encourages the dot product between matching image and text vectors (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle v_{i}\cdot w_{i}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle v_{i}\cdot w_{i}}</annotation>
</semantics>
</math></span><img src="./9a1ef9d1cb5b94b6dce3a0f0b6d6582391ddb2b7.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:6.07ex; height:2.009ex;" alt="{\displaystyle v_{i}\cdot w_{i}}" loading="lazy"></span>) to be high, while discouraging high dot products between non-matching pairs. The parameter <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle T>0}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>T</mi>
<mo>&gt;</mo>
<mn>0</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle T&gt;0}</annotation>
</semantics>
</math></span><img src="./e39e90b9e1e7d5be3b5eae57729dc63494bbe3fd.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:5.897ex; height:2.176ex;" alt="{\displaystyle T>0}" loading="lazy"></span> is the <a href="Temperature" title="Temperature">temperature</a>, which is parameterized in the original CLIP model as <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle T=e^{-\tau }}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>T</mi>
<mo>=</mo>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<mo>−<!-- − --></mo>
<mi>τ<!-- τ --></mi>
</mrow>
</msup>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle T=e^{-\tau }}</annotation>
</semantics>
</math></span><img src="./a44639f7aa4f80e75bd3e05a8a964570681924a7.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:8.179ex; height:2.509ex;" alt="{\displaystyle T=e^{-\tau }}" loading="lazy"></span> where <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \tau \in \mathbb {R} }">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>τ<!-- τ --></mi>
<mo>∈<!-- ∈ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="double-struck">R</mi>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \tau \in \mathbb {R} }</annotation>
</semantics>
</math></span><img src="./ae9dffeaac885108c5db356c1013f5d47fa22e11.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:5.721ex; height:2.176ex;" alt="{\displaystyle \tau \in \mathbb {R} }" loading="lazy"></span> is a learned parameter.
</p><p>Other loss functions are possible. For example, Sigmoid CLIP (SigLIP)<sup id="cite_ref-:4_6-0" class="reference"><a href="#cite_note-:4-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup> proposes the following loss function:<span class="mwe-math-element mwe-math-element-block"><span class="mwe-math-mathml-display mwe-math-mathml-a11y" style="display: none;"><math display="block" xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle L={\frac {1}{N}}\sum _{i,j\in 1:N}f((2\delta _{i,j}-1)(e^{\tau }w_{i}\cdot v_{j}+b))}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>L</mi>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mn>1</mn>
<mi>N</mi>
</mfrac>
</mrow>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo>∈<!-- ∈ --></mo>
<mn>1</mn>
<mo>:</mo>
<mi>N</mi>
</mrow>
</munder>
<mi>f</mi>
<mo stretchy="false">(</mo>
<mo stretchy="false">(</mo>
<mn>2</mn>
<msub>
<mi>δ<!-- δ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
</mrow>
</msub>
<mo>−<!-- − --></mo>
<mn>1</mn>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>τ<!-- τ --></mi>
</mrow>
</msup>
<msub>
<mi>w</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>v</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>j</mi>
</mrow>
</msub>
<mo>+</mo>
<mi>b</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle L={\frac {1}{N}}\sum _{i,j\in 1:N}f((2\delta _{i,j}-1)(e^{\tau }w_{i}\cdot v_{j}+b))}</annotation>
</semantics>
</math></span></span>where <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f(x)=\ln(1+e^{-x})}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>f</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mi>ln</mi>
<mo>⁡<!-- ⁡ --></mo>
<mo stretchy="false">(</mo>
<mn>1</mn>
<mo>+</mo>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<mo>−<!-- − --></mo>
<mi>x</mi>
</mrow>
</msup>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f(x)=\ln(1+e^{-x})}</annotation>
</semantics>
</math></span><img src="./8c86c2a2542256830d0512f51b55678717accd53.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:18.802ex; height:3.009ex;" alt="{\displaystyle f(x)=\ln(1+e^{-x})}" loading="lazy"></span> is the negative log <a href="Sigmoid_function" title="Sigmoid function">sigmoid</a> loss, and the <a href="Dirac_delta_function" title="Dirac delta function">Dirac delta symbol</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \delta _{i,j}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>δ<!-- δ --></mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \delta _{i,j}}</annotation>
</semantics>
</math></span><img src="./40e421f3ef0893f646c999fcc309a25ad6bad1f5.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -1.005ex; width:2.967ex; height:3.009ex;" alt="{\displaystyle \delta _{i,j}}" loading="lazy"></span> is 1 if <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle i=j}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>i</mi>
<mo>=</mo>
<mi>j</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle i=j}</annotation>
</semantics>
</math></span><img src="./706e0928b2bf0f24076b0c90bb20616ff2068343.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:4.859ex; height:2.509ex;" alt="{\displaystyle i=j}" loading="lazy"></span> else 0.
</p>
<div class="mw-heading mw-heading2"><h2 id="CLIP_models">CLIP models</h2></div>
<p>While the original model was developed by <a href="OpenAI" title="OpenAI">OpenAI</a>, subsequent models have been trained by other organizations as well.
</p>
<div class="mw-heading mw-heading3"><h3 id="Image_model">Image model</h3></div>

<p>The image encoding models used in CLIP are typically <a href="Vision_transformer" title="Vision transformer">vision transformers (ViT)</a>. The naming convention for these models often reflects the specific ViT architecture used. For instance, "ViT-L/14" means a "vision transformer large" (compared to other models in the same series) with a patch size of 14, meaning that the image is divided into 14-by-14 pixel patches before being processed by the transformer. The size indicator ranges from B, L, H, G (base, large, huge, giant), in that order.
</p><p>Other than ViT, the image model is typically a <a href="Convolutional_neural_network" title="Convolutional neural network">convolutional neural network</a>, such as <a href="Residual_neural_network" title="Residual neural network">ResNet</a> (in the original series by OpenAI), or ConvNeXt<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> (in the OpenCLIP model series by LAION<sup id="cite_ref-Ilharco_8-0" class="reference"><a href="#cite_note-Ilharco-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup>).
</p><p>Since the output vectors of the image model and the text model must have exactly the same length, both the image model and the text model have fixed-length vector outputs, which in the original report is called "embedding dimension".<sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>note 1<span class="cite-bracket">]</span></a></sup>
</p><p>For example, in the original OpenAI model, the ResNet models have embedding dimensions ranging from 512 to 1024,<sup id="cite_ref-:1_10-0" class="reference"><a href="#cite_note-:1-10"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Location: Table 19">: Table 19 </span></sup> and for the ViTs, from 512 to 768.<sup id="cite_ref-:1_10-1" class="reference"><a href="#cite_note-:1-10"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Location: Table 20">: Table 20 </span></sup>
</p>
<table class="wikitable">
<caption>Models released by OpenAI<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>note 2<span class="cite-bracket">]</span></a></sup>
</caption>
<tbody><tr>
<th>Model name
</th>
<th>Resolution
</th>
<th>Parameters (total, in millions)
</th>
<th>Parameters (vision)
</th>
<th>Parameters (text)
</th>
<th>Embedding dimension
</th>
<th>Size (MB)
</th>
<th>Release date
</th></tr>
<tr>
<td>RN50
</td>
<td>224
</td>
<td>102
</td>
<td>38.3
</td>
<td>63.1
</td>
<td>1024
</td>
<td>244
</td>
<td>2021-01
</td></tr>
<tr>
<td>RN101
</td>
<td>224
</td>
<td>120
</td>
<td>56.3
</td>
<td>63.1
</td>
<td>512
</td>
<td>278
</td>
<td>2021-03
</td></tr>
<tr>
<td>RN50x4
</td>
<td>288
</td>
<td>178
</td>
<td>87.1
</td>
<td>90.7
</td>
<td>640
</td>
<td>402
</td>
<td>2021-03
</td></tr>
<tr>
<td>RN50x16
</td>
<td>384
</td>
<td>291
</td>
<td>167.3
</td>
<td>123.0
</td>
<td>768
</td>
<td>630
</td>
<td>2021-07
</td></tr>
<tr>
<td>RN50x64
</td>
<td>448
</td>
<td>623
</td>
<td>420.4
</td>
<td>201.8
</td>
<td>1024
</td>
<td>1260
</td>
<td>2022-01
</td></tr>
<tr>
<td>ViT-B/32
</td>
<td>224
</td>
<td>151
</td>
<td>87.8
</td>
<td>63.1
</td>
<td>512
</td>
<td>338
</td>
<td>2021-01
</td></tr>
<tr>
<td>ViT-B/16
</td>
<td>224
</td>
<td>150
</td>
<td>86.2
</td>
<td>63.1
</td>
<td>512
</td>
<td>335
</td>
<td>2021-07
</td></tr>
<tr>
<td>ViT-L/14
</td>
<td>224
</td>
<td>428
</td>
<td>304.0
</td>
<td>123.0
</td>
<td>768
</td>
<td>890
</td>
<td>2022-01
</td></tr>
<tr>
<td>ViT-L/14@336px
</td>
<td>336
</td>
<td>428
</td>
<td>304.3
</td>
<td>123.0
</td>
<td>768
</td>
<td>891
</td>
<td>2022-04
</td></tr></tbody></table>
<p>Its implementation of ViT was the same as the original one,<sup id="cite_ref-:3_13-0" class="reference"><a href="#cite_note-:3-13"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup> with one modification: after position embeddings are added to the initial patch embeddings, there is a <a href="LayerNorm" class="mw-redirect" title="LayerNorm">LayerNorm</a>.
</p><p>Its implementation of <a href="Residual_neural_network" title="Residual neural network">ResNet</a> was the same as the original one,<sup id="cite_ref-resnet_14-0" class="reference"><a href="#cite_note-resnet-14"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> with 3 modifications:
</p>
<ul><li>In the start of the CNN (the "stem"), they used three stacked 3x3 convolutions instead of a single 7x7 convolution, as suggested by.<sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup></li>
<li>There is an average pooling of stride 2 at the start of each downsampling convolutional layer (they called it <i>rect-2 blur pooling</i> according to the terminology of <sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup>). This has the effect of blurring images before downsampling, for antialiasing.<sup id="cite_ref-17" class="reference"><a href="#cite_note-17"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup></li>
<li>The final convolutional layer is followed by a <a href="Pooling_layer#Vision_Transformer_pooling" title="Pooling layer">multiheaded attention pooling</a>.</li></ul>
<p>ALIGN a model with similar capabilities, trained by researchers from Google<sup id="cite_ref-:2_18-0" class="reference"><a href="#cite_note-:2-18"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> used <a href="EfficientNet" title="EfficientNet">EfficientNet</a>,<sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> a kind of <a href="Convolutional_neural_network" title="Convolutional neural network">convolutional neural network</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Text_model">Text model</h3></div>

<p>The text encoding models used in CLIP are typically <a href="Transformer_(deep_learning_architecture)" title="Transformer (deep learning architecture)">Transformers</a>.
</p><p>In the original OpenAI report, they reported using a Transformer (63M-parameter, 12-layer, 512-wide, 8 attention heads) with lower-cased <a href="Byte_pair_encoding" class="mw-redirect" title="Byte pair encoding">byte pair encoding</a> (BPE) with 49152 vocabulary size. Context length was capped at 76 for efficiency. Like <a href="Generative_pre-trained_transformer" title="Generative pre-trained transformer">GPT</a>, it was decoder-only, with only causally-masked self-attention.<sup id="cite_ref-:0_1-3" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap">: <span title="Page: 5
Quotation: &quot;Masked self-attention was used in the text encoder to preserve the ability to initialize with a pre-trained language model or add language modeling as an auxiliary objective, though exploration of this is left as future work.&quot;" class="tooltip tooltip-dashed" style="border-bottom: 1px dashed;">5</span> </sup> Its architecture is the same as <a href="GPT-2" title="GPT-2">GPT-2</a>.<sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup>
</p><p>Like <a href="BERT_(language_model)" title="BERT (language model)">BERT</a>, the text sequence is bracketed by two special tokens <code>[SOS]</code> and <code>[EOS]</code> ("start of sequence" and "end of sequence"). Take the activations of the highest layer of the transformer on the <code>[EOS]</code>, apply <a href="LayerNorm" class="mw-redirect" title="LayerNorm">LayerNorm</a>, then a final linear map. This is the text encoding of the input sequence. The final linear map has output dimension equal to the embedding dimension of whatever image encoder it is paired with. These models all had context length 77 and vocabulary size 49408.
</p><p>ALIGN<sup id="cite_ref-:2_18-1" class="reference"><a href="#cite_note-:2-18"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> used BERT of various sizes.
</p>
<div class="mw-heading mw-heading2"><h2 id="Dataset">Dataset</h2></div>
<div class="mw-heading mw-heading3"><h3 id="WebImageText">WebImageText</h3></div>
<p>The CLIP models released by OpenAI were trained on a dataset called "WebImageText" (WIT) containing 400 million pairs of images and their corresponding captions scraped from the internet. The total number of words in this dataset is similar in scale to the WebText dataset used for training <a href="GPT-2" title="GPT-2">GPT-2</a>, which contains about 40 gigabytes of text data.<sup id="cite_ref-:0_1-4" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p>The dataset contains 500,000 text-queries, with up to 20,000 (image, text) pairs per query. The text-queries were generated by starting with all words occurring at least 100 times in <a href="English_Wikipedia" title="English Wikipedia">English Wikipedia</a>, then extended by <a href="Bigram" title="Bigram">bigrams</a> with high <a href="Mutual_information" title="Mutual information">mutual information</a>, names of all Wikipedia articles above a certain search volume, and <a href="WordNet" title="WordNet">WordNet</a> <a href="Synset" title="Synset">synsets</a>.
</p><p>The dataset is private and has not been released to the public, and there is no further information on it.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>note 3<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="Data_preprocessing">Data preprocessing</h4></div>
<p>For the CLIP image models, the input images are preprocessed by first dividing each of the R, G, B values of an image by the maximum possible value, so that these values fall between 0 and 1, then subtracting by <code>[0.48145466, 0.4578275, 0.40821073]</code>, and dividing by <code>[0.26862954, 0.26130258, 0.27577711]</code>.
</p><p>The rationale was that these are the mean and standard deviations of the images in the WebImageText dataset, so this preprocessing step roughly <a href="Whitening_transformation" title="Whitening transformation">whitens</a> the image tensor. These numbers slightly differ from the standard preprocessing for ImageNet, which uses <code>[0.485, 0.456, 0.406]</code> and <code>[0.229, 0.224, 0.225]</code>.<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup>
</p><p>If the input image does not have the same resolution as the native resolution (224×224 for all except ViT-L/14@336px, which has 336×336 resolution), then the input image is scaled down by <a href="Bicubic_interpolation" title="Bicubic interpolation">bicubic interpolation</a>, so that its shorter side is the same as the native resolution, then the central square of the image is <a href="Cropping_(image)" title="Cropping (image)">cropped</a> out.
</p>
<div class="mw-heading mw-heading3"><h3 id="Others">Others</h3></div>
<p>ALIGN<sup id="cite_ref-:2_18-2" class="reference"><a href="#cite_note-:2-18"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> used over one billion image-text pairs, obtained by extracting images and their <a href="Alt_attribute" title="Alt attribute">alt-tags</a> from online crawling. The method was described as similar to how the Conceptual Captions dataset<sup id="cite_ref-24" class="reference"><a href="#cite_note-24"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup> was constructed, but instead of complex filtering, they only applied a frequency-based filtering.
</p><p>Later models trained by other organizations had published datasets. For example, <a href="LAION" title="LAION">LAION</a> trained OpenCLIP with published datasets LAION-400M, LAION-2B, and DataComp-1B.<sup id="cite_ref-25" class="reference"><a href="#cite_note-25"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Ilharco_8-1" class="reference"><a href="#cite_note-Ilharco-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Training">Training</h2></div>
<p>In the original OpenAI CLIP report, they reported training 5 <a href="Residual_neural_network" title="Residual neural network">ResNet</a> and 3 <a href="Vision_transformer" title="Vision transformer">ViT</a> (ViT-B/32, ViT-B/16, ViT-L/14). Each was trained for 32 epochs. The largest ResNet model took 18 days to train on 592 <a href="Volta_(microarchitecture)#Products" title="Volta (microarchitecture)">V100</a> GPUs. The largest ViT model took 12 days on 256 V100 GPUs.
</p><p>All ViT models were trained on 224×224 image resolution. The ViT-L/14 was then boosted to 336×336 resolution by FixRes,<sup id="cite_ref-26" class="reference"><a href="#cite_note-26"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup> resulting in a model.<sup id="cite_ref-27" class="reference"><a href="#cite_note-27"><span class="cite-bracket">[</span>note 4<span class="cite-bracket">]</span></a></sup> They found this was the best-performing model.<sup id="cite_ref-:0_1-5" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup class="reference nowrap"><span title="Location: Appendix F. Model Hyperparameters">: Appendix F. Model Hyperparameters </span></sup>
</p><p>In the OpenCLIP series, the ViT-L/14 model was trained on 384 <a href="Ampere_(microarchitecture)" title="Ampere (microarchitecture)">A100</a> GPUs on the LAION-2B dataset, for 160 epochs for a total of 32B samples seen.<sup id="cite_ref-28" class="reference"><a href="#cite_note-28"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Applications">Applications</h2></div>
<div class="mw-heading mw-heading3"><h3 id="Cross-modal_retrieval">Cross-modal retrieval</h3></div>
<p>CLIP's cross-modal retrieval enables the alignment of visual and textual data in a shared latent space, allowing users to retrieve images based on text descriptions and vice versa, without the need for explicit image annotations.<sup id="cite_ref-29" class="reference"><a href="#cite_note-29"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup> In <b>text-to-image retrieval</b>, users input descriptive text, and CLIP retrieves images with matching embeddings. In <b>image-to-text retrieval</b>, images are used to find related text content.
</p><p>CLIP’s ability to connect visual and textual data has found applications in multimedia search, content discovery, and recommendation systems.<sup id="cite_ref-30" class="reference"><a href="#cite_note-30"><span class="cite-bracket">[</span>26<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-31" class="reference"><a href="#cite_note-31"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Image_classification">Image classification</h3></div>
<p>CLIP can perform zero-shot image classification tasks. This is achieved by prompting the text encoder with class names and selecting the class whose embedding is closest to the image embedding. For example, to classify an image, they compared the embedding of the image with the embedding of the text "A photo of a {class}.", and the {class} that results in the highest dot product is outputted.
</p>
<div class="mw-heading mw-heading3"><h3 id="CLIP_for_multimodal_learning">CLIP for multimodal learning</h3></div>
<p>CLIP has been used as a component in <a href="Multimodal_learning" title="Multimodal learning">multimodal learning</a>. For example, during the training of <a href="Google_DeepMind" title="Google DeepMind">Google DeepMind</a>'s Flamingo (2022),<sup id="cite_ref-32" class="reference"><a href="#cite_note-32"><span class="cite-bracket">[</span>28<span class="cite-bracket">]</span></a></sup> the authors trained a CLIP pair, with BERT as the text encoder and NormalizerFree ResNet F6<sup id="cite_ref-33" class="reference"><a href="#cite_note-33"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup> as the image encoder. The image encoder of the CLIP pair was taken with parameters frozen and the text encoder was discarded. The frozen image encoder was then combined with a frozen <a href="Chinchilla_(language_model)" title="Chinchilla (language model)">Chinchilla language model</a>, by finetuning with some further parameters that connect the two frozen models.
</p>
<div class="mw-heading mw-heading3"><h3 id="Applications_in_other_domains">Applications in other domains</h3></div>
<ul><li>CLIP's image encoder is a pre-trained image <a href="Feature_learning" title="Feature learning">featurizer</a>. This can then be fed into other AI models.<sup id="cite_ref-:0_1-6" class="reference"><a href="#cite_note-:0-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>Models like <a href="Stable_Diffusion" title="Stable Diffusion">Stable Diffusion</a> use CLIP's text encoder to transform text prompts into embeddings for image generation.<sup id="cite_ref-stable-diffusion-github_3-1" class="reference"><a href="#cite_note-stable-diffusion-github-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> CLIP can also be used as a gradient signal for directly guiding diffusion ("CLIP guidance")<sup id="cite_ref-34" class="reference"><a href="#cite_note-34"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-35" class="reference"><a href="#cite_note-35"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup> or other generative art.<sup id="cite_ref-36" class="reference"><a href="#cite_note-36"><span class="cite-bracket">[</span>32<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>Fine-tuned CLIP models can be used to rank images by aesthetic quality, which may be useful as a step in filtering a large dataset into a smaller one with higher quality.<sup id="cite_ref-37" class="reference"><a href="#cite_note-37"><span class="cite-bracket">[</span>33<span class="cite-bracket">]</span></a></sup></li></ul>
<ul><li>CLIP can be used to <a href="Image_captioning" class="mw-redirect" title="Image captioning">generate image captions</a> by matching text inputs to image embeddings.<sup id="cite_ref-38" class="reference"><a href="#cite_note-38"><span class="cite-bracket">[</span>34<span class="cite-bracket">]</span></a></sup></li></ul>
<div class="mw-heading mw-heading2"><h2 id="Notes">Notes</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */


.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}


/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap"><ol class="references">
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text">Similar to the "embedding dimension" of text embedding in Transformer models.</span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text"><div class="mw-highlight mw-highlight-lang-python mw-content-ltr" dir="ltr"><pre><span class="err">!</span><span class="n">pip</span> <span class="n">install</span> <span class="n">git</span><span class="o">+</span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">openai</span><span class="o">/</span><span class="n">CLIP</span><span class="o">.</span><span class="n">git</span>
<span class="err">!</span><span class="n">wget</span> <span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">openai</span><span class="o">/</span><span class="n">CLIP</span><span class="o">/</span><span class="n">raw</span><span class="o">/</span><span class="n">main</span><span class="o">/</span><span class="n">CLIP</span><span class="o">.</span><span class="n">png</span> <span class="o">-</span><span class="n">O</span> <span class="n">CLIP</span><span class="o">.</span><span class="n">png</span>

<span class="kn">import</span><span class="w"> </span><span class="nn">torch</span>
<span class="kn">import</span><span class="w"> </span><span class="nn">clip</span>
<span class="kn">from</span><span class="w"> </span><span class="nn">PIL</span><span class="w"> </span><span class="kn">import</span> <span class="n">Image</span>
<span class="kn">import</span><span class="w"> </span><span class="nn">numpy</span><span class="w"> </span><span class="k">as</span><span class="w"> </span><span class="nn">np</span>

<span class="n">device</span> <span class="o">=</span> <span class="s2">"cuda"</span> <span class="k">if</span> <span class="n">torch</span><span class="o">.</span><span class="n">cuda</span><span class="o">.</span><span class="n">is_available</span><span class="p">()</span> <span class="k">else</span> <span class="s2">"cpu"</span>
<span class="k">for</span> <span class="n">m</span> <span class="ow">in</span> <span class="n">clip</span><span class="o">.</span><span class="n">available_models</span><span class="p">():</span>
<span class="n">model</span><span class="p">,</span> <span class="n">preprocess</span> <span class="o">=</span> <span class="n">clip</span><span class="o">.</span><span class="n">load</span><span class="p">(</span><span class="n">m</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
<span class="n">input_resolution</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">visual</span><span class="o">.</span><span class="n">input_resolution</span>
<span class="n">context_length</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">context_length</span>
<span class="n">vocab_size</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">vocab_size</span>

<span class="nb">print</span><span class="p">(</span><span class="s2">"Model parameters:"</span><span class="p">,</span> <span class="sa">f</span><span class="s2">"</span><span class="si">{</span><span class="n">np</span><span class="o">.</span><span class="n">sum</span><span class="p">([</span><span class="nb">int</span><span class="p">(</span><span class="n">np</span><span class="o">.</span><span class="n">prod</span><span class="p">(</span><span class="n">p</span><span class="o">.</span><span class="n">shape</span><span class="p">))</span><span class="w"> </span><span class="k">for</span><span class="w"> </span><span class="n">p</span><span class="w"> </span><span class="ow">in</span><span class="w"> </span><span class="n">model</span><span class="o">.</span><span class="n">parameters</span><span class="p">()])</span><span class="si">:</span><span class="s2">,</span><span class="si">}</span><span class="s2">"</span><span class="p">)</span>
<span class="nb">print</span><span class="p">(</span><span class="s2">"Input resolution:"</span><span class="p">,</span> <span class="n">input_resolution</span><span class="p">)</span>
<span class="nb">print</span><span class="p">(</span><span class="s2">"Context length:"</span><span class="p">,</span> <span class="n">context_length</span><span class="p">)</span>
<span class="nb">print</span><span class="p">(</span><span class="s2">"Vocab size:"</span><span class="p">,</span> <span class="n">vocab_size</span><span class="p">)</span>

<span class="n">n_params_vision</span> <span class="o">=</span> <span class="nb">sum</span><span class="p">(</span><span class="n">p</span><span class="o">.</span><span class="n">numel</span><span class="p">()</span> <span class="k">for</span> <span class="n">p</span> <span class="ow">in</span> <span class="n">model</span><span class="o">.</span><span class="n">visual</span><span class="o">.</span><span class="n">parameters</span><span class="p">())</span>
<span class="n">n_params_text</span> <span class="o">=</span> <span class="nb">sum</span><span class="p">(</span><span class="n">p</span><span class="o">.</span><span class="n">numel</span><span class="p">()</span> <span class="k">for</span> <span class="n">p</span> <span class="ow">in</span> <span class="n">model</span><span class="o">.</span><span class="n">transformer</span><span class="o">.</span><span class="n">parameters</span><span class="p">())</span>
<span class="n">image</span> <span class="o">=</span> <span class="n">preprocess</span><span class="p">(</span><span class="n">Image</span><span class="o">.</span><span class="n">open</span><span class="p">(</span><span class="s2">"CLIP.png"</span><span class="p">))</span><span class="o">.</span><span class="n">unsqueeze</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span><span class="o">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
<span class="n">image_features</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">encode_image</span><span class="p">(</span><span class="n">image</span><span class="p">)</span>
<span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s2">"Model: </span><span class="si">{</span><span class="n">m</span><span class="si">}</span><span class="s2">, #vision parameters: </span><span class="si">{</span><span class="n">n_params_vision</span><span class="si">:</span><span class="s2">,</span><span class="si">}</span><span class="s2">, #text parameters: </span><span class="si">{</span><span class="n">n_params_text</span><span class="si">:</span><span class="s2">,</span><span class="si">}</span><span class="s2">, embedding dimension: </span><span class="si">{</span><span class="n">image_features</span><span class="o">.</span><span class="n">shape</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span><span class="si">}</span><span class="s2">"</span><span class="p">)</span>
<span class="k">del</span> <span class="n">model</span><span class="p">,</span> <span class="n">preprocess</span><span class="p">,</span> <span class="n">image</span><span class="p">,</span> <span class="n">image_features</span>
</pre></div></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text">It is not the same as the Wikipedia-based Image Text dataset, also called "WIT".<sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup></span>
</li>
<li id="cite_note-27"><span class="mw-cite-backlink"><b><a href="#cite_ref-27">^</a></b></span> <span class="reference-text">They referred to this as both <code>ViT-L/14-336px</code> and <code>ViT-L/14@336px</code>, inconsistently throughout the report.</span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<div class="reflist reflist-columns references-column-width" style="column-width: 30em;">
<ol class="references">
<li id="cite_note-:0-1"><span class="mw-cite-backlink">^ <a href="#cite_ref-:0_1-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:0_1-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-:0_1-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-:0_1-3"><sup><i><b>d</b></i></sup></a> <a href="#cite_ref-:0_1-4"><sup><i><b>e</b></i></sup></a> <a href="#cite_ref-:0_1-5"><sup><i><b>f</b></i></sup></a> <a href="#cite_ref-:0_1-6"><sup><i><b>g</b></i></sup></a></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */


.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}


/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFRadfordKimHallacyRamesh2021" class="citation conference cs1">Radford, Alec; Kim, Jong Wook; Hallacy, Chris; Ramesh, Aditya; Goh, Gabriel; Agarwal, Sandhini; Sastry, Girish; Askell, Amanda; Mishkin, Pamela; Clark, Jack; Krueger, Gretchen; Sutskever, Ilya (2021-07-01). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v139/radford21a"><i>Learning Transferable Visual Models From Natural Language Supervision</i></a>. Proceedings of the 38th International Conference on Machine Learning. PMLR. pp.&nbsp;<span class="nowrap">8748–</span>8763.</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite id="CITEREFHendriksenBleekerVakulenkovan_Noord2022" class="citation journal cs1">Hendriksen, Mariya; Bleeker, Maurits; Vakulenko, Svitlana; van Noord, Nanne; Kuiper, Ernst; de Rijke, Maarten (2022). Hagen, Matthias; Verberne, Suzan; Macdonald, Craig; Seifert, Christin; Balog, Krisztian; Nørvåg, Kjetil; Setty, Vinay (eds.). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://link.springer.com/chapter/10.1007/978-3-030-99736-6_20">"Extending CLIP for Category-to-Image Retrieval in E-Commerce"</a></span>. <i>Advances in Information Retrieval</i>. Cham: Springer International Publishing: <span class="nowrap">289–</span>303. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2F978-3-030-99736-6_20">10.1007/978-3-030-99736-6_20</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-3-030-99736-6</bdi>.</cite></span>
</li>
<li id="cite_note-stable-diffusion-github-3"><span class="mw-cite-backlink">^ <a href="#cite_ref-stable-diffusion-github_3-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-stable-diffusion-github_3-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://github.com/CompVis/stable-diffusion">"Stable Diffusion Repository on GitHub"</a>. CompVis - Machine Vision and Learning Research Group, LMU Munich. 17 September 2022. <a rel="nofollow" class="external text" href="https://web.archive.org/web/20230118183342/https://github.com/CompVis/stable-diffusion">Archived</a> from the original on January 18, 2023<span class="reference-accessdate">. Retrieved <span class="nowrap">17 September</span> 2022</span>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text"><cite class="citation cs2"><a rel="nofollow" class="external text" href="https://github.com/LAION-AI/aesthetic-predictor"><i>LAION-AI/aesthetic-predictor</i></a>, LAION AI, 2024-09-06<span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-08</span></span></cite></span>
</li>
<li id="cite_note-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-5">^</a></b></span> <span class="reference-text"><cite id="CITEREFSohn2016" class="citation journal cs1">Sohn, Kihyuk (2016). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper/2016/hash/6b180037abbebea991d8b1232f8a8ca9-Abstract.html">"Improved Deep Metric Learning with Multi-class N-pair Loss Objective"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>29</b>. Curran Associates, Inc.</cite></span>
</li>
<li id="cite_note-:4-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-:4_6-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFZhaiMustafaKolesnikovBeyer2023" class="citation conference cs1">Zhai, Xiaohua; Mustafa, Basil; Kolesnikov, Alexander; Beyer, Lucas (2023). <a rel="nofollow" class="external text" href="https://openaccess.thecvf.com/content/ICCV2023/html/Zhai_Sigmoid_Loss_for_Language_Image_Pre-Training_ICCV_2023_paper.html"><i>Sigmoid Loss for Language Image Pre-Training</i></a>. IEEE/CVF International Conference on Computer Vision (ICCV). pp.&nbsp;<span class="nowrap">11975–</span>11986.</cite></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text"><cite id="CITEREFLiuMaoWuFeichtenhofer2022" class="citation conference cs1">Liu, Zhuang; Mao, Hanzi; Wu, Chao-Yuan; Feichtenhofer, Christoph; Darrell, Trevor; Xie, Saining (2022). <a rel="nofollow" class="external text" href="https://openaccess.thecvf.com/content/CVPR2022/html/Liu_A_ConvNet_for_the_2020s_CVPR_2022_paper.html"><i>A ConvNet for the 2020s</i></a>. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp.&nbsp;<span class="nowrap">11976–</span>11986.</cite></span>
</li>
<li id="cite_note-Ilharco-8"><span class="mw-cite-backlink">^ <a href="#cite_ref-Ilharco_8-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Ilharco_8-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFIlharcoWortsmanWightmanGordon2021" class="citation cs2">Ilharco, Gabriel; Wortsman, Mitchell; Wightman, Ross; Gordon, Cade; <a href="Nicholas_Carlini" title="Nicholas Carlini">Carlini, Nicholas</a>; Taori, Rohan; Dave, Achal; Shankar, Vaishaal; Namkoong, Hongseok (July 2021), <a rel="nofollow" class="external text" href="https://github.com/mlfoundations/open_clip"><i>OpenCLIP</i></a>, <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.5281%2Fzenodo.5143773">10.5281/zenodo.5143773</a><span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-06</span></span></cite></span>
</li>
<li id="cite_note-:1-10"><span class="mw-cite-backlink">^ <a href="#cite_ref-:1_10-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:1_10-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFRadfordKimHallacyRamesh2021" class="citation arxiv cs1">Radford, Alec; Kim, Jong Wook; Hallacy, Chris; Ramesh, Aditya; Goh, Gabriel; Agarwal, Sandhini; Sastry, Girish; Askell, Amanda; Mishkin, Pamela; Clark, Jack; Krueger, Gretchen; Sutskever, Ilya (2021). "Learning Transferable Visual Models From Natural Language Supervision". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2103.00020">2103.00020</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text"><cite class="citation cs2"><a rel="nofollow" class="external text" href="https://github.com/openai/CLIP/"><i>openai/CLIP</i></a>, OpenAI, 2024-09-06<span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-06</span></span></cite></span>
</li>
<li id="cite_note-:3-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-:3_13-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFDosovitskiyBeyerKolesnikovWeissenborn2021" class="citation arxiv cs1">Dosovitskiy, Alexey; Beyer, Lucas; Kolesnikov, Alexander; Weissenborn, Dirk; Zhai, Xiaohua; Unterthiner, Thomas; Dehghani, Mostafa; Minderer, Matthias; Heigold, Georg; Gelly, Sylvain; Uszkoreit, Jakob (2021-06-03). "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2010.11929">2010.11929</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-resnet-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-resnet_14-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFHeZhangRenSun2015" class="citation conference cs1">He, Kaiming; Zhang, Xiangyu; Ren, Shaoqing; Sun, Jian (10 Dec 2015). <i>Deep Residual Learning for Image Recognition</i>. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1512.03385">1512.03385</a></span>.</cite></span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-15">^</a></b></span> <span class="reference-text"><cite id="CITEREFHeZhangZhangZhang2018" class="citation arxiv cs1">He, Tong; Zhang, Zhi; Zhang, Hang; Zhang, Zhongyue; Xie, Junyuan; Li, Mu (2018-12-05). "Bag of Tricks for Image Classification with Convolutional Neural Networks". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1812.01187">1812.01187</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text"><cite id="CITEREFZhang2018" class="citation journal cs1">Zhang, Richard (2018-09-27). <a rel="nofollow" class="external text" href="https://openreview.net/forum?id=SklVEnR5K7">"Making Convolutional Networks Shift-Invariant Again"</a>.</cite> <span class="cs1-visible-error citation-comment"><code class="cs1-code">{{cite journal}}</code>: </span><span class="cs1-visible-error citation-comment">Cite journal requires <code class="cs1-code">|journal=</code> (help)</span></span>
</li>
<li id="cite_note-17"><span class="mw-cite-backlink"><b><a href="#cite_ref-17">^</a></b></span> <span class="reference-text"><cite id="CITEREFZhang2019" class="citation arxiv cs1">Zhang, Richard (2019-06-08). "Making Convolutional Networks Shift-Invariant Again". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1904.11486">1904.11486</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-:2-18"><span class="mw-cite-backlink">^ <a href="#cite_ref-:2_18-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-:2_18-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-:2_18-2"><sup><i><b>c</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFJiaYangXiaChen2021" class="citation journal cs1">Jia, Chao; Yang, Yinfei; Xia, Ye; Chen, Yi-Ting; Parekh, Zarana; Pham, Hieu; Le, Quoc; Sung, Yun-Hsuan; Li, Zhen; Duerig, Tom (2021-07-01). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v139/jia21b.html">"Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision"</a>. <i>Proceedings of the 38th International Conference on Machine Learning</i>. PMLR: <span class="nowrap">4904–</span>4916.</cite></span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text"><cite id="CITEREFTanLe2020" class="citation arxiv cs1">Tan, Mingxing; Le, Quoc V. (2020-09-11). "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1905.11946">1905.11946</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.LG">cs.LG</a>].</cite></span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text"><cite id="CITEREFRadfordWuChildLuan2019" class="citation journal cs1">Radford, Alec; Wu, Jeff; Child, R.; Luan, D.; Amodei, Dario; Sutskever, I. (2019). "Language Models are Unsupervised Multitask Learners". <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:160025533">160025533</a>.</cite> <span class="cs1-visible-error citation-comment"><code class="cs1-code">{{cite journal}}</code>: </span><span class="cs1-visible-error citation-comment">Cite journal requires <code class="cs1-code">|journal=</code> (help)</span></span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text"><cite id="CITEREFSrinivasanRamanChenBendersky2021" class="citation book cs1">Srinivasan, Krishna; Raman, Karthik; Chen, Jiecao; Bendersky, Michael; Najork, Marc (2021-07-11). "WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning". <i>Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</i>. pp.&nbsp;<span class="nowrap">2443–</span>2449. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2103.01913">2103.01913</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1145%2F3404835.3463257">10.1145/3404835.3463257</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-1-4503-8037-9</bdi>.</cite></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://github.com/openai/CLIP/issues/20">"std and mean for image normalization different from ImageNet · Issue #20 · openai/CLIP"</a>. <i>GitHub</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2024-09-19</span></span>.</cite></span>
</li>
<li id="cite_note-24"><span class="mw-cite-backlink"><b><a href="#cite_ref-24">^</a></b></span> <span class="reference-text"><cite id="CITEREFSharmaDingGoodmanSoricut2018" class="citation journal cs1">Sharma, Piyush; Ding, Nan; Goodman, Sebastian; Soricut, Radu (July 2018). Gurevych, Iryna; Miyao, Yusuke (eds.). <a rel="nofollow" class="external text" href="https://aclanthology.org/P18-1238/">"Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning"</a>. <i>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</i>. Melbourne, Australia: Association for Computational Linguistics: <span class="nowrap">2556–</span>2565. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.18653%2Fv1%2FP18-1238">10.18653/v1/P18-1238</a></span>.</cite></span>
</li>
<li id="cite_note-25"><span class="mw-cite-backlink"><b><a href="#cite_ref-25">^</a></b></span> <span class="reference-text"><cite id="CITEREFChertiBeaumontWightmanWortsman2023" class="citation book cs1">Cherti, Mehdi; Beaumont, Romain; Wightman, Ross; Wortsman, Mitchell; Ilharco, Gabriel; Gordon, Cade; Schuhmann, Christoph; Schmidt, Ludwig; Jitsev, Jenia (June 2023). "Reproducible Scaling Laws for Contrastive Language-Image Learning". <i>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</i>. pp.&nbsp;<span class="nowrap">2818–</span>2829. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2212.07143">2212.07143</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FCVPR52729.2023.00276">10.1109/CVPR52729.2023.00276</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>979-8-3503-0129-8</bdi>.</cite></span>
</li>
<li id="cite_note-26"><span class="mw-cite-backlink"><b><a href="#cite_ref-26">^</a></b></span> <span class="reference-text"><cite id="CITEREFTouvronVedaldiDouzeJegou2019" class="citation journal cs1">Touvron, Hugo; Vedaldi, Andrea; Douze, Matthijs; Jegou, Herve (2019). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper/2019/hash/d03a857a23b5285736c4d55e0bb067c8-Abstract.html">"Fixing the train-test resolution discrepancy"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>32</b>. Curran Associates, Inc.</cite></span>
</li>
<li id="cite_note-28"><span class="mw-cite-backlink"><b><a href="#cite_ref-28">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://huggingface.co/laion/CLIP-ViT-L-14-laion2B-s32B-b82K">"laion/CLIP-ViT-L-14-laion2B-s32B-b82K · Hugging Face"</a>. <i>huggingface.co</i>. 2023-09-10<span class="reference-accessdate">. Retrieved <span class="nowrap">2024-09-06</span></span>.</cite></span>
</li>
<li id="cite_note-29"><span class="mw-cite-backlink"><b><a href="#cite_ref-29">^</a></b></span> <span class="reference-text"><cite id="CITEREFHendriksenBleekerVakulenkovan_Noord2021" class="citation arxiv cs1">Hendriksen, Mariya; Bleeker, Maurits; Vakulenko, Svitlana; van Noord, Nanne; Kuiper, Ernst; de Rijke, Maarten (2021). "Extending CLIP for Category-to-image Retrieval in E-commerce". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2112.11294">2112.11294</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-30"><span class="mw-cite-backlink"><b><a href="#cite_ref-30">^</a></b></span> <span class="reference-text"><cite id="CITEREFBeaumont2024" class="citation cs2">Beaumont, Romain (2024-09-07), <a rel="nofollow" class="external text" href="https://github.com/rom1504/clip-retrieval"><i>rom1504/clip-retrieval</i></a><span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-08</span></span></cite></span>
</li>
<li id="cite_note-31"><span class="mw-cite-backlink"><b><a href="#cite_ref-31">^</a></b></span> <span class="reference-text"><cite id="CITEREFHaltakov2024" class="citation cs2">Haltakov, Vladimir (2024-09-03), <a rel="nofollow" class="external text" href="https://github.com/haltakov/natural-language-image-search"><i>haltakov/natural-language-image-search</i></a><span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-06</span></span></cite></span>
</li>
<li id="cite_note-32"><span class="mw-cite-backlink"><b><a href="#cite_ref-32">^</a></b></span> <span class="reference-text"><cite id="CITEREFAlayracDonahueLucMiech2022" class="citation journal cs1">Alayrac, Jean-Baptiste; Donahue, Jeff; Luc, Pauline; Miech, Antoine; Barr, Iain; Hasson, Yana; Lenc, Karel; Mensch, Arthur; Millican, Katherine; Reynolds, Malcolm; Ring, Roman; Rutherford, Eliza; Cabi, Serkan; Han, Tengda; Gong, Zhitao (2022-12-06). <a rel="nofollow" class="external text" href="https://proceedings.neurips.cc/paper_files/paper/2022/hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html">"Flamingo: a Visual Language Model for Few-Shot Learning"</a>. <i>Advances in Neural Information Processing Systems</i>. <b>35</b>: <span class="nowrap">23716–</span>23736.</cite></span>
</li>
<li id="cite_note-33"><span class="mw-cite-backlink"><b><a href="#cite_ref-33">^</a></b></span> <span class="reference-text"><cite id="CITEREFBrockDeSmithSimonyan2021" class="citation journal cs1">Brock, Andy; De, Soham; Smith, Samuel L.; Simonyan, Karen (2021-07-01). <a rel="nofollow" class="external text" href="https://proceedings.mlr.press/v139/brock21a.html">"High-Performance Large-Scale Image Recognition Without Normalization"</a>. <i>Proceedings of the 38th International Conference on Machine Learning</i>. PMLR: <span class="nowrap">1059–</span>1071.</cite></span>
</li>
<li id="cite_note-34"><span class="mw-cite-backlink"><b><a href="#cite_ref-34">^</a></b></span> <span class="reference-text"><cite id="CITEREFRameshDhariwalNicholChu2022" class="citation arxiv cs1">Ramesh, Aditya; Dhariwal, Prafulla; Nichol, Alex; Chu, Casey; Chen, Mark (2022-04-12). "Hierarchical Text-Conditional Image Generation with CLIP Latents". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2204.06125">2204.06125</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
<li id="cite_note-35"><span class="mw-cite-backlink"><b><a href="#cite_ref-35">^</a></b></span> <span class="reference-text"><cite id="CITEREFHendriksenBleekerVakulenkovan_Noord2022" class="citation journal cs1">Hendriksen, Mariya; Bleeker, Maurits; Vakulenko, Svitlana; van Noord, Nanne; Kuiper, Ernst; de Rijke, Maarten (2022). Hagen, Matthias; Verberne, Suzan; Macdonald, Craig; Seifert, Christin; Balog, Krisztian; Nørvåg, Kjetil; Setty, Vinay (eds.). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://link.springer.com/chapter/10.1007/978-3-030-99736-6_20">"Extending CLIP for Category-to-Image Retrieval in E-Commerce"</a></span>. <i>Advances in Information Retrieval</i>. Cham: Springer International Publishing: <span class="nowrap">289–</span>303. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1007%2F978-3-030-99736-6_20">10.1007/978-3-030-99736-6_20</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-3-030-99736-6</bdi>.</cite></span>
</li>
<li id="cite_note-36"><span class="mw-cite-backlink"><b><a href="#cite_ref-36">^</a></b></span> <span class="reference-text"><cite id="CITEREFWhitaker2022" class="citation web cs1">Whitaker, Jonathan (2022-05-22). <a rel="nofollow" class="external text" href="https://wandb.ai/johnowhitaker/nca/reports/Fun-with-Neural-Cellular-Automata--VmlldzoyMDQ5Mjg0">"Fun With Neural Cellular Automata"</a>. <i>W&amp;B</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2024-09-08</span></span>.</cite></span>
</li>
<li id="cite_note-37"><span class="mw-cite-backlink"><b><a href="#cite_ref-37">^</a></b></span> <span class="reference-text"><cite class="citation cs2"><a rel="nofollow" class="external text" href="https://github.com/LAION-AI/aesthetic-predictor"><i>LAION-AI/aesthetic-predictor</i></a>, LAION AI, 2024-09-06<span class="reference-accessdate">, retrieved <span class="nowrap">2024-09-08</span></span></cite></span>
</li>
<li id="cite_note-38"><span class="mw-cite-backlink"><b><a href="#cite_ref-38">^</a></b></span> <span class="reference-text"><cite id="CITEREFMokadyHertzBermano2021" class="citation arxiv cs1">Mokady, Ron; Hertz, Amir; Bermano, Amit H. (2021). "ClipCap: CLIP Prefix for Image Captioning". <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2111.09734">2111.09734</a></span> [<a rel="nofollow" class="external text" href="https://arxiv.org/archive/cs.CV">cs.CV</a>].</cite></span>
</li>
</ol></div>
<div class="mw-heading mw-heading2"><h2 id="External_links">External links</h2></div>
<ul><li><a rel="nofollow" class="external text" href="https://openai.com/research/clip">OpenAI's CLIP webpage</a></li>
<li><a rel="nofollow" class="external text" href="https://github.com/mlfoundations/open_clip">OpenCLIP: An open source implementation of CLIP</a></li>
<li><cite id="CITEREFArora2023" class="citation web cs1">Arora, Aman (2023-03-11). <a rel="nofollow" class="external text" href="https://amaarora.github.io/posts/2023-03-11_Understanding_CLIP_part_2.html">"The Annotated CLIP (Part-2)"</a>. <i>amaarora.github.io</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2024-09-11</span></span>.</cite></li></ul></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-06-21" href="https://en.wikipedia.org/wiki/?title=Contrastive_Language-Image_Pre-training&amp;oldid=1296674956">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>

</body></html>